Nature Biotechnology
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match Nature Biotechnology's content profile, based on 172 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.
Nilsson, A.; Sporre, E.; Schulte, D.; Snijder, J.; Edfors, F.; Käll, L.
Show abstract
Reading a proteins sequence from tandem mass spectra without a reference is limited by single-spectrum accuracy, most acutely across the hypervariable complementaritydetermining regions of antibodies. Broadly specific proteases tile a protein with long, overlapping peptides, so every residue is covered by many independent de novo reads. borgonovo assembles their per-step probability profiles into a reference-free per-residue consensus, seeding templates from mass-closure-consistent reads and recruiting the rest by substitution- tolerant alignment and per-column voting. Re-decoding each spectrum with a prior from its consensus position lifts amino acid accuracy on placed spectra from 0.80 to 0.87. On the therapeutic antibody trastuzumab, nine proteases cover its heavy and light chains completely at 0.88 fixed-window identity, and 0.93 on the pruned assembly once local indels are accommodated. Applied unchanged to five secretome proteins and trastuzumab with three proteases, it reaches 0.87 mean fixed-window identity over 82% coverage. borgonovo is open source and works with most de novo sequencers, so redundant digestion turns any of them into a protein sequencer where no reference exists.
Goode, Z.; Tiedemann, E.; Ben Ameur, L.; Pavan, K.; Young, K.; Sek, M.; Nevue, A.; Zhu, J.; Houghton, J.; Fu, Y.; Boisvert, H.; Saunders, A.
Show abstract
Probe-based genomics technologies are extending molecular analysis into intact tissues and fixed cells, yet strategies to decode complex experimental conditions encoded in cellular RNA remain limited. Here we present a modular framework that integrates custom software tools with purpose-built cloning reagents to design, assemble, validate, and deploy combinatorial DNA barcodes. Combinatorial barcodes comprise spatially adjacent collections of known sequences, enabling millions of unique molecules to be efficiently distinguished using a limited set of probes. Our software tools integrate with optimized assembly plasmids and whole plasmid long-read sequencing for high-fidelity construction and structural validation of diverse combinatorial barcode architectures. Assembled barcode libraries are flexibly transferred into user-modified expression vectors to support diverse downstream experimental applications. We showcase the versatility of this framework by assembling two structurally distinct combinatorial barcode libraries, each containing millions of unique sequences. Following rabies virus-based delivery to the mouse brain, we validate in vivo decoding of a combinatorial barcode architecture capable of distinguishing ~16.3 million expressed RNAs through probe-based in situ sequencing. Our framework for flexible and accurate combinatorial barcode construction fills a technically demanding niche delivering cost-effective molecular reagents for multiplexed experimentation on current and evolving probe-based genomics platforms.
Robles-Remacho, A.; Zou, Y.; Jensen, A.; Tricopoulos, C.; Grillo, M.; Nilsson, M.
Show abstract
The spatial organization of post-transcriptional regulation is a fundamental yet difficult to access layer of tissue biology. MicroRNAs (miRNAs) are small RNAs with a key role in post-transcriptional regulation, but their short length has excluded them from spatial profiling technologies, leaving them largely unexplored in spatial transcriptomics. Here, we introduce miR-Space, a method that converts individual miRNAs into extended, uniquely barcoded molecules directly in tissue, enabling their spatial detection by in situ sequencing. Across 30 mouse and human brain sections, miR-Space enabled highly multiplexed miRNA profiling at single-molecule and single-cell resolution, joint analysis with mRNA, and implementation on the automated Xenium platform. miR-Space resolved major anatomical regions and cell populations from spatial miRNA expression, identified reproducible cell-associated miRNA signatures, and uncovered previously unknown spatial and cellular distributions of multiple miRNAs. Together, these capabilities establish miR-Space as a framework for integrating miRNAs into spatial transcriptomics, enabling spatial miRNomics at anatomical and single-cell resolution.
Zhang, H.; Wang, P.; Zhao, Y.; Yang, L.; Xue, T.; Liu, L.; Zhao, Y.; Zhang, Z.; Ma, J.; Zeng, B.; Zhang, P.; Wang, C.; Pan, D.; Gao, Z.; Liu, Z.; Zeng, Z.
Show abstract
Spatial transcriptomics offers a glimpse into the immunology of tissues. However, limitations in spatial transcriptomics preclude the detection of highly diverse, low-abundance, and previously unknown sequences, including immune repertoires and microbiota. Here, we introduce Archimap, a spatial transcriptomic platform that simultaneously profiles spatial transcriptomes, immune repertoires, and microbiota from formalin-fixed paraffin-embedded (FFPE) tissues. Using Archimap, we profile the spatial localization of TCRs, BCRs, and the microbiota landscape in archived clinical tissues at single-cell resolution. Through comprehensive benchmarking, we validate Archimaps performance and fidelity. Archimap in situ assembles the immune complex and reconstructs the clonal evolution of antibodies. Together, Archimap shows the power of in situ discovery of functional immune repertoires for their antitumor immunity.
Birk, S.; Merchant, A.; Vahidi, A.; Theis, F. J.; Lotfollahi, M.
Show abstract
Spatially-resolved transcriptomics (SRT) measures gene expression at single-cell resolution while preserving each cells spatial location, enabling the joint study of cell identity and cellular niche, the recurring microenvironment that organises tissue function. Existing representation-learning methods typically capture only one of these axes at a time. We present SQUINT, a graph vector-quantized variational autoencoder (VQ-VAE) that learns two disjoint codebooks per cell from a shared architecture: a cell codebook quantising the per-cell embedding before neighbourhood aggregation, biased toward cell-intrinsic identity, and a niche codebook quantising the embedding after graph neural network (GNN) aggregation, biased toward spatial context. Both use residual vector quantization, giving a coarse-to-fine discrete-token hierarchy. SQUINT is trained with per-branch negative-binomial reconstruction objectives and three domain-motivated components that we show are crucial: a within-section cosine adjacency loss that anchors the niche codes in the spatial graph, a cross-section contrastive loss on the cell latents that aligns transcriptomically matched cells, and a decoder section covariate that absorbs batch effects. Across three datasets spanning four spatial assays (STARmap, MERFISH, CosMx, Xenium) and four tasks - niche identification, cell-type identification, cross-section integration, and spatial gene-expression imputation in held-out regions - SQUINT outperforms or is competitive with strong baselines on identification and achieves the most faithful cross-section integration. The resulting discrete vocabulary makes tissues directly consumable by transformer-style foundation models and enables one-step query-to-reference atlas mapping via code-distribution similarity, which we demonstrate on a CosMx human non-small-cell lung cancer cohort.
Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.
Show abstract
Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.
Jamhuri, M.; Irawan, A.
Show abstract
Alignment-free lineage assignment from k-mer frequency profiles is widely used for SARS-CoV-2 surveillance, and the methods that do it are ranked against each other by margins of one or two percentage points. Those rankings rest on an unchecked protocol. Public repositories hold many near-duplicate genomes, and stratified random splitting puts members of such a group on both sides of the split, so a classifier is credited for sequences it has already seen. We propose quantised profile hashing, which finds near duplicates in k-mer feature space by rounding each frequency vector and hashing it. No sequence is compared with any other, so one pass over the feature matrix suffices and no similarity threshold has to be chosen. Rounding is also what makes the groups well defined, and they are then kept whole across the training, validation and test sets. On 255,611 genomes from seven Pango lineages, random splitting leaves 5.09% of test sequences with a near duplicate in training, on a benchmark ranked by margins of one or two points. Ten update rules were trained twice, identically except for the partition. The contaminated benchmark separates one rule from the leader at 0.05; the clean one separates none. The two orderings are uncorrelated, Kendall{tau} = +0.022, with rules moving 3.2 positions on average and the leader of one benchmark ranking eighth on the other. A ranking obtained under contamination therefore says nothing about the ranking without it, and the quantity worth reporting beside a score is the leakage rate of the split.
shen, x.; ZHANG, X.-Y.
Show abstract
Joint single-cell transcriptomic-metabolomic profiling remains technically intractable. Here we present CHIMERA (Cell-level Hybrid Inference of Metabolome Embedded on RNA Atlas), a data-driven framework that learns transcriptome-to-metabolome mappings from spatially paired multi-omics data and transfers them to unpaired scRNA-seq. CHIMERA generates quantitative, database-independent single-cell metabolite abundances and, by pairing them with the measured transcriptome of the same cells, enables joint co-embedding of genes and metabolites for the discovery of differential metabolites and co-regulated gene-metabolite modules. Using 10x Visium paired with MALDI-MSI from murine liver sections and a matched scRNA-seq reference, CHIMERA achieves a per-metabolite median Pearson r = 0.285 with positive cross-section generalization. On an independent Liver Cell Atlas Western-diet cohort, CHIMERA recovers metabolic reprogramming that recapitulate published non-alcoholic fatty liver disease pathophysiology. Applied to a Rarres2 (chemerin) knock-down hepatocellular carcinoma model, CHIMERA uncovers metabolic heterogeneity among tumour-associated macrophages, resolving four metabolic subclusters (MC-0 to MC-3); Rarres2 appears to drive macrophage polarization from an LAM-like MC-3 state toward Spp1+ like MC-0/MC-2 by modulating a co-regulated gene-metabolite module--a dual-omics phenotype undetectable by either modality alone. CHIMERA is the first data-driven framework for quantitative single-cell metabolome inference, opening joint transcriptomic- metabolomic analyses inaccessible to either experimental or knowledge-based computational approaches.
Verma, S.; Arora, N.; Ajay, C. P.; Singh, P.; Mallick, H.; Ghosh, T. S.
Show abstract
Deciphering gut microbiome to host metabolome interaction is critical for understanding how microbial communities generate bioactive signals that shape host physiology and disease. Progress, however, has been hindered by inconsistent metabolite annotations, poor interoperability across studies, and the absence of integrated resources placing microbiome-derived metabolites within their functional, microbial, physiological, and clinical context. Here we present HuMMANet (Human Microbiome Metabolome Annotation Network), a harmonized resource integrating 46 paired gut microbiome metabolome studies (59 study-units; 14,405 samples; 13 disease categories plus a healthy/control reference category) with a scalable metabolite-harmonization framework. HuMMANet resolves heterogeneous annotations through a multi-stage workflow spanning RefMet, HMDB, PubChem, Metabolomics Workbench, SMPDB, MiMeDB 2.0, GNPS/microbeMASST, DrugBank, and DrugCentral, yielding a reference atlas of 54,914 unique metabolites, annotated with standardized chemical identifiers, biochemical pathways, microbial producer associations, physiological distributions, disease links, and structural relationships to approved therapeutics, a unified reference framework for microbiome metabolome research. Applying HuMMANet to a multi-cohort integration of adult serum and fecal metabolomes, we identified 519 serum and 322 fecal metabolites reproducibly associated with gut microbial community composition (PERMANOVA, P < 0.05 in at least 50% of studies in which detected), enriched for specific biomolecular classes and pathways. Cross-referencing these against Health Associated Core Keystone (HACK) taxa revealed 58 serum and 25 fecal metabolites (HACK positive) whose taxon-level associations tracked positively with the taxon specific HACK indices. These reproducible metabolomic signatures of microbiome health included indole3propionic acid, a gut barrier-protective microbial tryptophan metabolite, and 3phenylpropionate. Drug similarity annotation within HuMMANet linked 16 of this serum and 13 fecal HACK positive metabolites to therapeutics used in neurological, inflammatory, and vascular disease. Conversely, 38 serum and 65 fecal metabolites, including imidazole propionate and long-chain acylcarnitines such as ACar 18:0, showed HACK negative signatures previously associated with dysbiosis-linked disease. GNPS/microbeMASST and MiMeDB 2.0 annotations further traced subsets of these metabolites to putative bacterial producers. HuMMANet thus provides a standardized framework for reproducible microbiome metabolome integration, enabling cross study discovery and translational prioritization of conserved microbiome derived metabolic signatures across human populations and disease states.
Zhu, W.; Tian, M.; Duan, Y.; Reisman, S. J.; Miller, S. E.; Corden, E.; ter Weele, M.; Song, L.; Blount, J.; Safi, A.; Schreiber, J.; Gersbach, C. A.; Crawford, G. E.; Gordan, R.
Show abstract
CRISPR technologies based on nuclease-deactivated Cas9 (dCas9) rely on programmable DNA binding rather than DNA cleavage, yet the intrinsic DNA-recognition properties that govern optimal guide RNA (gRNA) performance remain poorly understood. Existing approaches either measure genomic occupancy in cells or infer dCas9 behavior from cleavage-based Cas9 datasets, despite DNA binding being substantially more permissive than DNA cleavage. Here we introduce TANGO (Targeted Array-based Nucleic acid-Guided Occupancy), a high-density DNA-array platform that quantitatively profiles intrinsic dCas9:gRNA binding across tens of thousands of DNA targets in a cell-free system. TANGO captures established features of dCas9 target recognition, while providing substantially greater sensitivity than prior assays. Comparison with ChIP-seq data demonstrates that intrinsic DNA-binding specificity is a major driver of genomic occupancy and reveals that chromatin accessibility modulates the intrinsic binding affinity required for dCas9 recruitment. Across CRISPRi/a guides, TANGO identifies multiple independent biochemical determinants of guide performance--including on-target affinity, mismatch tolerance, and ribonucleoprotein assembly--and flags problematic and highly promiscuous guides overlooked by current specificity metrics. Unexpectedly, some guides retain substantial guide-directed DNA binding even in the absence of a protospacer-adjacent motif (PAM), revealing an additional dimension of dCas9 specificity. Together, these results establish intrinsic DNA recognition as a quantitative and experimentally accessible determinant of dCas9 function, providing a framework for improving guide selection and enhancing the precision of CRISPR technologies.
Wang, Y. V.; Park, J.; Kim, M. C.; Mazumder, T.; Sonpal, K.; Bikaran, M.; Steinhart, Z.; Schmidt, R.; Sun, Y.; Lee, S.-H.; Marson, A.; Ye, C. J.; Hwang, B.
Show abstract
Surface proteins define T cell identity and function, but the abundance of each protein is not determined by transcription alone. Existing genome-wide CRISPR screens in primary human T cells either profile the transcriptome or isolate cells based on a single functional or protein phenotype. Here we present SCITO-Perturb-seq, a novel platform that couples combinatorial-indexed single-cell cytometry sequencing with pooled CRISPR activation (CRISPRa) to map the causal regulation of 201 surface proteins across 3.6 million human CD4 T cells. We find that 16% of activated genes significantly alter the expression of at least one surface protein. By applying semi-nonnegative matrix factorization to the perturbation effect matrix, we identified five modules corresponding to known CD4 T cell states. Notably, these modules group surface proteins by their shared response to perturbation, revealing coordinated regulation of proteins that are not co-expressed in unperturbed cells. SCITO-Perturb-seq represents the first genome-wide CRISPRa screen paired with direct, high-dimensional surface protein profiling, providing a comprehensive regulatory map of the CD4 T cell surface proteome.
Soitu, C.; Sahin, U.; Magnussen, A.; Wong, A.; Bonnaffe, W.; Fan, M.; Bilici, M.; Davis, S.; Fischer, R.; Reese, J.; House, T.; Moradi, S.; Teague, R.; Ansorge, O.; Malacrino, S.; Alham, N. K.; McGregor, E.; Maldonado-Perez, D.; Tomlinson, I.; Wedge, D.; Hester, J.; Issa, F.; Edwards, C.; Bryant, R.; Mills, I.; Rittscher, J.; Hamdy, F.; Woodcock, D.; Verrill, C.; Rao, S.
Show abstract
Spatially resolved DNA sequencing holds promise due to its potential utility in understanding cancer intra-tumour heterogeneity and tumour evolution in relation to tissue architecture. However, it has so far been used to a limited extent due to technical challenges and high cost of existing methods. Hence we aimed to develop a high throughput spatial genomic assay to obtain copy number alteration (CNA) information at user-defined spatial resolution. We derived CNA profiles from ultra-low coverage whole genome sequencing at sub-millimetre resolution from archival samples using a novel method called Adaptive Resolution Multiscale Spatial DNA sequencing (ARMS DNAseq). We used it to profile CNAs from more than 766 regions (tiles) from 3 patients, covering a total area of over 300 mm2, with 1.2-2.6 million mapped reads per tile and tile sizes of 0.1-0.99mm2. Using ARMS DNAseq, we delineate tumour evolution in a spatial context, and identify more tumour subclones that were obscured or incompletely represented in bulk multi-region whole genome sequencing. Next, we show associations between tumour subclones and morphology, and prediction of subclone identity from deep learning-derived image representations. Finally, we demonstrate multi-omic integration by alignment with spatial transcriptomic data, showing subclone-specific immune cell co-occurrence as well as transcriptional programmes cutting across subclone boundaries. ARMS DNAseq converts low-throughput, region-by-region profiling into a scalable and adaptable workflow for direct spatial copy number profiling from archival tissue sections.
Lee, J.; Glazier, J.; Weichselbaum, R. R.; Mimee, M.
Show abstract
Engineered bacteria offer a distinct modality for cancer therapy by exploiting the ability of certain species to colonize tumors and deliver therapeutic payloads. Improving their efficacy and safety requires control over bacterial activity after tumor colonization, yet few microbial chassis permit it. Bifidobacterium longum, a probiotic with intrinsic tumor-targeting and antitumor activity, is a promising chassis but lacks such control. Here, we develop a genetic control system that regulates B. longum activity within tumors, from gene expression to bacterial abundance. A human-isolate-derived replicon supports plasmid maintenance without antibiotic selection, and promoter and ribosome-binding-site libraries provide [~]150-fold and [~]48-fold expression ranges, respectively. Signal peptides enable secretion of structurally diverse therapeutic payloads and B. longum secreting CCL21 or an anti-PD-L1 nanobody reduces tumor growth relative to PBS controls. Anhydrotetracycline delivered in drinking water induces transgene expression in tumor-resident bacteria and reduces intratumoral bacterial load through CRISPRi targeting essential genes. Together, these results establish a tumor-homing probiotic as an externally controllable therapeutic chassis.
Wang, S.; Zhu, B.; Li, S.; Wei, X.
Show abstract
High-resolution sequencing-based spatial transcriptomics, including Stereo-seq and Visium HD, aggregates dense capture units into cell-resolved expression matrices. During tissue processing and permeabilization, RNA released from source cells can spread to neighbouring capture locations, reducing cell-type specificity and biasing downstream analyses. Here we developed SPARKLE (Spatial Ambient RNA Kernel-based Leakage Estimator), a cell-level correction method that uses capture locations outside cell-segmentation masks as within-sample spatial evidence of leakage. SPARKLE fits sparse spatial kernels to out-of-mask observations to estimate a sample-level propagation scale and gene-specific leakage coefficients. It corrects only genes supported by out-of-mask goodness of fit and uses expression-dependent conservative shrinkage to protect highly expressing source cells. Across ten simulated scenarios, SPARKLE achieved the highest cell-wise concordance and the lowest RMSE in 9 of 10 scenarios. In axolotl brain, mouse brain and human ovarian cancer, SPARKLE removed ectopic marker signal from neighbouring cells while retaining source-cell expression, improved agreement with independent single-cell and single-nucleus references, and recovered COL1A2-SDC4 signaling of fibroblast origin that collagen diffusion had obscured. Conclusions remained stable across plausible spatial scales and background-bin sizes. Runtime scaled near-linearly with tissue-window area, and was further acceleration on GPU. SPARKLE is therefore a reference-free, fast and scalable method for correcting local RNA leakage from evidence contained within each sample, improving the reliability of cell-type localization, tissue-compartment identification and cell-cell communication inference.
Shi, T. H.; Sinclair, J. A.; Gao, F.; Senapati, S.; Moorman, T.; Chang, H.-C.
Show abstract
Viral diagnostics during early phases of infection are often limited by target scarcity and the deployment tempo. We significantly advance both quantitative accuracy and diagnostic throughput of viral agglutination assays with Immuno-Janus Particle (IJP) aggregation behavior that "flicker" stochastically with size-dependent statistics. By scrutinizing microscale blinking patterns of time series fluorescent videos, we decipher Brownian dynamics of individual IJP-Virus conjugates and IJP aggregates via windowed Ito stochastic analysis (termed the Culsans method). High-frequency rotational fluctuation is deconvolved from corrupting drifts caused by gravitational sedimentation and Brownian translational motion. This methodology enables a non-linear mapping of angular positions of detected IJPs and IJP aggregates to extract rotational diffusivity (Dr) (and subsequently overall construct size) with superior linearity (R2[≥]0.85). The aggregation behavior exhibits a maximum when the IJP and viral particle concentrations are equal. The virion-bridged IJP-IJP conjugates significantly shift the detectable hydrodynamic diameter in the Poisson limit of reduced virus concentration with respect to IJPs, pushing the limit of detection (LOD) to 103 - 104 virions per mL in untreated human plasma. This tunable platform offers a rapid, low-volume, and scalable alternative to lab-based RT-PCR, bridging the gap between virion sensitivity and field-readiness.
Salaudeen, A. L.; Shyiak, T.; de Boer, C. G.
Show abstract
Virus-like particles (VLPs) enable transient, non-integrating delivery of CRISPR-Cas9 ribonucleoprotein cargo. Although VLPs have been reported for efficient DNA editing via base editors RNP delivery, the diversity of base editors tested as VLPs remains limited. We generated and benchmarked a panel of 12 base editors on the v5 eVLP backbone, targeting three genomic loci (HEK3, B2M, PDCD1) across five VLP dosages in LentiX-293T cells. Editing efficiency was generally dosage-dependent across all editors and varied by editor class and identity; PAM-flexible variants had lower editing efficiency than NGG-restricted counterparts, and the dual-function SPACE base editors showed reduced efficiency. We further characterized position-specific editing efficiencies and outcomes of the base editor VLP collection, revealing that a wide variety of mutation types are possible with the base editors in this collection.
Wu, S.; Mahajan, M.; van der Linde, R. M.; Zhu, C.; van IJzendoorn, D.; West, R. B.; Matusiak, M.
Show abstract
Single-cell spatial transcriptomics is now central to studying tumors in their native tissue context. Here we present the first comprehensive, independent evaluation of Atera, a new spatial whole transcriptome platform, compared against Xenium on adjacent sections of human ductal carcinoma in situ (DCIS). We show that Atera enables granular cell-state annotation and resolves rare cell populations, which we experimentally validate by multiplex immunofluorescence (IF). We further show that its transcriptome-wide coverage enables inference of copy-number alterations at single-cell resolution, allowing us to reconstruct the clonal evolution of DCIS. We orthogonally confirm the inferred copy-number alterations by whole-genome sequencing of 16 microdissected tumor regions from a consecutive tissue section. Finally, by mapping the immune microenvironment onto this clonal architecture, we demonstrate the feasibility of tracking the changes in immune response along the clonal tumor evolution in situ. Together, our results establish Atera as a validated platform for tracking clonal evolution and immune adaptation in clinical samples.
Yagi, S.; Sagami, N.; Eshima, I.; Hiramatsu, K.
Show abstract
Label-free Raman imaging of living cells is photon limited: at exposures compatible with cellular dynamics, single-pixel spectra carry about one count per channel on a dominant smooth background. We present an unmixing framework in which the decoder of a physics-constrained autoencoder is restricted to a data-driven spectroscopic dictionary: band centers,widths, and pseudo-Voigt shapes are measured from the dataset and fixed, and the network learns only nonnegative band amplitudes, a smooth B-spline background, and a per-pixel gain.First, on slit-scanning images of HeLa cells (532 nm) the dictionary yields spike-free component spectra that read as band tables, including a resonance-enhanced cytochrome-c-associated component matching literature spectra, and the most stable decomposition against the component number. Second, the dictionary and initialization calibrated at 1 s exposure perline transfer to 100 ms per line (12 s sweeps): cytochrome-c spectral identity survives a single sweep (correlation 0.92) while its map remains photon limited; the dictionary provides spectral physicality, and the transferred initialization prevents a structural collapse that global map correlations miss; in a measurement-derived phantom the dictionary estimator holds thecytochrome-c spectrum to 17-19{degrees} spectral angle at 100 ms, where classical factorizations and free decoders lose it (55-64{degrees}). Estimation on the count-equivalent detector output uses a calibrated shifted-Poisson quasi-likelihood. Third, evaluation must be time matched:correlation against a separately acquired reference saturates through slow specimen drift and acquisition mismatch rather than photon noise, and the self-consistency of learned denoisers is inflated by shared bias; time-matched self-consistency and independent cross-checks areproposed.
Gencturk, M. M.; Cicek, A. E.
Show abstract
Single-cell RNA sequencing (scRNA-seq) is widely used to infer copy number profiles from tumor cells. Existing methods build on a reference-based normalization paradigm: normalizing each tumor cell against a reference of normal cells, whether supplied, in-sample, or synthesized. This makes them reference-dependent and as a result, sensitive to cohort composition, and prone to false positives. To address these limitations, we introduce REFCON, a deep-learning model that enables reference-free copy number profiling from scRNA-seq data. REFCON estimates local copy-number deviations and jointly optimizes them into a genome-wide per-cell profile. It profiles pure tumors, generalizes to unseen tissues and platforms, and stays robust to cohort composition. Predicted copy number profiles distinguish malignant cells with high specificity, producing far fewer false-positive calls, and improve clonal reconstruction. The model can also benefit from reference cells when available, turning a field requirement into an optional refinement. Hence, REFCON extends reliable per-cell copy number profiling to the scRNA-seq data collected without matched normals.
Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.
Show abstract
Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.